Skip to content

laya: add a reproduction command for every Apple Silicon benchmark number - #67

Open
cacheline999 wants to merge 4 commits into
ThinkFlowLab:mainfrom
cacheline999:laya-review-followups
Open

cacheline999 wants to merge 4 commits into
ThinkFlowLab:mainfrom
cacheline999:laya-review-followups

Conversation

@cacheline999

@cacheline999 cacheline999 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Follow-ups to #30:

  1. Every performance number in the Apple Silicon recipe now has a command that reproduces it. Laya on Apple Silicon: worker, benchmarks and recipe #30
    shipped scripts for the main comparisons, but several numbers in the recipe came from one-off probes
    that were not in the repository: requests after an idle gap, the cost of new input lengths and the
    memory they add, a checkpoint loaded while serving, and the CPU fallback. The benchmark README now
    has a table from each recipe claim to the command behind it, and the missing pieces are added:

    • paired.py --gap SECONDS (each request after that much idle, on a fresh connection) and --only;
    • lengths.py: first request of new input lengths and the footprint they add, for one or two flag sets;
    • late_load.py: a checkpoint loaded while serving, its first request directly and through the frontend;
    • fallback.py: Laya's fallback to the CPU under a lowered MPS memory limit;
    • release.py: the memory torch.mps.empty_cache() gives back after many new lengths, and whether
      those lengths are cold again afterwards. It prepares the model with the worker's own startup
      (build_app: options and warmup), exits if the model is not on MPS, walks lengths with
      lengths.py's walk, and reads each footprint --settle seconds after a release;
    • a loop of fresh starts for the first request after ready, alternating the worker with plain
      laya-serve.

    late_load.py, fallback.py and release.py refuse a measured run on battery or under load like
    the other scripts (--feasibility runs anyway); the check and the workload reader are now shared
    helpers in env.py.

    Two statements had no command behind them and were changed instead: the recipe no longer gives
    numbers for keeping the GPU busy between requests (it only says the worker does not do this), and
    the ~70 s late load with --compile could not be reproduced (19–22 s alone, 17 s with a second
    worker holding memory on the same GPU), so the recipe gives no number for a load past the frontend's
    60 s and only says what happens then. The recipe's count of every length up to the window is now the
    454 that the full-window command reaches (it said 477, from an earlier probe with shorter inputs).

    Each is an A/B where the claim is a comparison (with the options against without, or a worker against
    plain laya-serve). A rerun gives other numbers on another Mac or under other load; the comparison is
    what should carry over. Rerunning everything on the M1 Pro changed these recipe statements (see Test
    Result): the late load with --compile (19–22 s; no number past 60 s), the CPU fallback request
    (30–73 s), every length up to the window (454 lengths, 5.2 GB with the options, 3.7 GB without),
    what a release gives back (back to about 2.95 GB), and plain laya-serve's first request after ready
    (0.7–1.1 s then, 0.2–0.4 s in seven later fresh starts, cause not known).

  2. Readiness contract test (flaky on an M4 in the last review): it now checks that the first request
    after ready succeeds with a valid answer, without a latency bound. That latency depends on how long
    the GPU has been idle and on the Mac (an M5 answered a first request without any warmup in 132–143 ms),
    so it stays in the benchmarks. That the warmup ran before ready is still tested.

  3. ruff format drift: the Laya files are formatted with ruff 0.16.10 defaults (88 columns, as the rest
    of the repository; the first commit is formatting only, with every file's syntax tree unchanged), and
    requirements-mps.txt pins ruff, pytest and httpx2. Comments that the formatter had pushed onto
    closing brackets are back above their statements, so the two noqa: SIM115 suppress again.

  4. Stale docstring in tests/laya/test_worker.py; the reviewer's M4 is listed in the recipe.

No change to the worker's runtime code. A CI job for tests/laya was suggested in the review; happy to
add one in a separate PR if you want it.

Test Plan

System1-Omni Version / Commit: 62dc52a on top of 58b8cbe

  • Every reproduction command in the benchmark README run at least once on the M1 Pro.
  • ruff format --check and ruff check --select E4,E7,E9,F on the Laya files, ruff 0.16.10.
  • PYTHONPATH=src python -m pytest tests/laya, and with LAYA_CONTRACT=1; contract tests on MPS
    without and with --compile --weights fp16.
  • The documented-commands test also checks the inline commands of the reproduction table against each
    script's flags (a misspelled flag fails it).
  • mkdocs build --strict.

Test Result

All reruns on the M1 Pro (16 GB). None was on an idle machine on mains power, so they were run as
feasibility; what each shows is that the command runs and whether the recipe's effect is there. The
conditions column says which session a row comes from.

claim recipe in #30 rerun conditions
request after 2 s idle, options / none ~0.9 (105–115 vs 114–127 ms) 0.89 (950 vs 1020 ms, 30 pairs, wide interval) battery, load 9–17
first request of a new length, extra 15 ms with the options, 6 without 28 and 13 ms (25 and 8 ms over all 454); 16–18 ms with the options in release.py battery, load 9–17; AC, load 6–9
footprint per 100 new lengths +520 MB with the options, +64 without +539 and +72 MB battery, load 9–17
footprint after every length up to the window 5.3 GB with the options, 4.0 GB without (477 lengths) 5.2 GB (from 2.8) and 3.7 GB (from 3.8) over 454 lengths → recipe updated battery, load 9–17
empty_cache() after 100 new lengths, with the options releases most of it; those lengths pay again footprint after a release 2946–2953 MB in three runs (start 2942–3016, after the lengths 3231–3481), 2896–2897 after running them again and releasing again → recipe: back to about 2.95 GB AC, load 6–9
checkpoint loaded while serving, first request 5–10 s without, ~70 s with --compile 6.5 and 8.5 s without; 18.6–22 s with (200 through the frontend); 17 s with a second worker holding GPU memory → recipe: 19–22 s, no number past 60 s battery, load 9–17 (6.5, 19–22 s); AC, load 6–8 (8.5, 18.6 s); battery, load 5 (17 s)
CPU fallback, the request that ran out of memory ~30 s 73 s plain, 68 s with the options, 33 s on a second plain run, 31 and 45 s later → recipe: 30–73 s battery and AC, load 5–17
CPU requests after the fallback 140–270 ms 148–273 ms (the first one 1–14 s) as above
first request after ready, worker with the options 70–81 ms 67 and 75 ms AC, under load
first request after ready, plain laya-serve 0.7–1.1 s 227 ms; then 206–270 ms right after a worker run and 213–362 ms after 30 s idle (three each) → recipe gives both ranges, cause not known AC, load 15; battery, load 4–6
checkpoint download 97 s 401 s (network) —

Checks: ruff format and check clean; 98 unit tests passed, 118 with LAYA_CONTRACT=1; contract on MPS
20/20 plain and with the options; strict docs build passes.

Self-review

Several self-review passes ran before this description; each issue below was reproduced (a test that
failed, or a rerun) before it was fixed.

  • lengths.py now skips the lengths the warmup ran (read back from the worker) and stops at the window;
    the full-window walk covers 454 lengths, 57–512 tokens.

  • late_load.py waits up to --timeout, follows up a load cut off by a 504 or a timeout only if the
    checkpoint became resident, refuses a --late the worker already serves, and documents every field.

  • release.py prepares the model the way the worker does, exits off MPS, keeps the traceback of other
    failures, reads W1 by id, and walks lengths with lengths.py.

  • Every script that measures refuses a noisy machine; fallback.py says why its memory probe failed.

  • Two measurements did not hold and are not in the recipe: a release giving back "about half" and then
    "about 60%" came from a starting footprint that varies by a few hundred MB between processes and from
    reading the footprint before the release had landed (release.py now reads --settle seconds later);
    a torch.mps.synchronize() before the release made no difference in an A/B (it waits 0.004 ms after
    a request). The ~70 s late load could not be reproduced, with or without a second worker.

  • I have reviewed the full diff and addressed the issues I found.

  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.

  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.

  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

…0 defaults

Formatting only (88 columns, as the rest of the repository); no behaviour change.
… M4 in the recipe

- The contract test of the first request after ready asserts that it succeeds with a valid answer. Its
  latency depends on how long the GPU has been idle and on the Mac (an M5 answered a first request without
  warmup in 132-143 ms), so it is left to the benchmarks; that the warmup ran before ready is tested
  separately.
- requirements-mps.txt pins pytest, httpx2 and ruff; the recipe's Test section runs the ruff checks.
- The recipe lists the reviewer's M4 among the Macs the tests ran on.
- Test docstring points at tests/laya.
@cacheline999 cacheline999 changed the title Laya: follow-ups from the #30 review laya: add a reproduction command for every Apple Silicon benchmark number Oct 3, 2026
…cipe

The benchmark README now has a table from each recipe claim to the command behind it. New
scripts cover what ThinkFlowLab#30 measured with one-off probes:

- paired.py --gap SECONDS (each request after that much idle, on a fresh connection) and --only;
- lengths.py: the first request of new input lengths and the footprint they add, for one or two
  flag sets; it skips the lengths the warmup ran and stops at the window;
- late_load.py: a checkpoint loaded while serving, directly and through the frontend;
- fallback.py: Laya's fallback to the CPU under a lowered MPS memory limit;
- release.py: in one process prepared by the worker's build_app, what torch.mps.empty_cache()
  gives back after new lengths. It releases once before the walk and reads every footprint
  --settle seconds after a release, because the footprint shows a release up to ~2 s late and
  the starting footprint varies by a few hundred MB between runs; where a release lands does not;
- a loop of fresh starts alternating the worker with plain laya-serve.

Every measuring script refuses a run on battery or above --max-load (late_load, fallback and
release unless --feasibility); env.py holds the shared check and workload reader.

Recipe changes from the reruns: a late load with --compile took 19-22 s (the ~70 s in ThinkFlowLab#30 was
not reproduced, so the recipe gives no number for a load past the frontend's 60 s); the CPU
fallback request took 30-73 s; every length up to the window is 454 lengths, 5.2 GB with the
options and 3.7 GB without; a release brings the footprint back to about 2.95 GB. The heartbeat
numbers had no script and are gone.

Comments that ruff format had pushed onto closing brackets are back above their statements
(syntax trees unchanged). No change to the worker's runtime code.
…ince

The recipe gave 0.7-1.1 s for plain laya-serve's first request after ready, from three runs on
2026-09-28. Seven fresh starts since, with the same torch, laya and macOS, took 0.2-0.4 s, both
right after another run and after 30 s of idle, so the GPU's state left by a previous run does
not explain it. The recipe now gives both and says the cause is not known. The worker's own
first request (67-81 ms) is unchanged.

The README says every script that measures refuses a noisy machine; report.py does not measure.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant